Acta Crystallographica Section D Structural Biology
● International Union of Crystallography (IUCr)
Preprints posted in the last 30 days, ranked by how well they match Acta Crystallographica Section D Structural Biology's content profile, based on 59 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Burton-Smith, R. N.; Murata, K.
Show abstract
Here, we present MC-Bayes, a Python-based script for processing cryo-electron microscopy EER movies on one or more GPUs using MotionCor3 in a user-friendly manner. Further, it generates the .star files necessary for RELION to perform Bayesian polishing (a.k.a.: reference-based motion correction) with EER movies. Until now, Bayesian polishing of EER data was only possible if the CPU-based "RELIONCor" implementation of MotionCor2 was used, which is sub-optimal on GPU-heavy cryo-EM processing systems. This wrapper was created for those facilities and/or users who may have (many) powerful GPUs, but for whatever reason have few CPU cores or less system RAM. Leveraging MotionCor3, MC-Bayes allows motion correction of EER data 2 or more times faster (depending on system) than the RELION CPU implementation, except in circumstances where dozens or hundreds of CPU cores with high quantities of system RAM can be utilised.
Panjikar, S.; Weiss, M.; Jayatilaka, D.
Show abstract
Directional anisotropy in electron density provides key information about chemical bonding that is not readily accessible from conventional electron-density maps. Here, a model-independent framework is presented for decomposing experimental structure factors into angular components using spherical harmonics. Reciprocal-space projection onto spherical harmonics followed by standard Fourier synthesis yields angularly filtered density maps. The{ell} = 0 component captures the isotropic part of the density, while the{ell} = 1 components resemble px, py and pz-like dipolar functions that highlight directional electronic structure. Applications to high-resolution datasets, including urea, the Gly-Ala dipeptide and a 0.97 [A]{beta}-lactamase structure, reveal chemically interpretable dipolar features associated with carbonyl and amide bonds, N-H interactions and aromatic{pi} systems. Quantitative analysis using bond-centred sampling demonstrates stable dipolar signatures that remain detectable under moderate resolution truncation. These results establish spherical-harmonic angular decomposition as a practical framework for extracting directional electronic information from crystallographic electron-density maps. SynopsisAngular decomposition of experimental structure factors reveals dipolar anisotropy and directional electron-density features that are directly meaningful for chemical interpretation.
Miyaguchi, I.; Hata, H.; Kuribayashi, T.; Takahashi, S.; Kashima, A.; Murasaki, K.; Matsumoto, S.; Terayama, K.; Ohta, M.; Ikeguchi, M.
Show abstract
Accurate assessment of ligand coordinate-density consistency across different resolutions remains challenging in macromolecular crystallography. We introduce the atomic Box Correlation Coefficient (aBCC), an atom-level metric for evaluating the consistency between ligand atomic coordinates and electron density in a resolution-standardized framework. To predict aBCC values from electron-density maps, we developed QAEmap, a machine-learning model based on three-dimensional convolutional neural networks (3D-CNNs). The model was trained using Fourier-truncated electron-density maps and corresponding ligand coordinates generated from high-resolution structures in the Protein Data Bank. It was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures. was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures.The prediction accuracy gradually decreased with decreasing resolution, but remained reliable up to [~]3.5 [A]. These results demonstrate that aBCC enables resolution-standardized atom-wise evaluation of coordinate-density consistency across different resolutions and provide a foundation for further development and refinement of machine learning-based coordinate validation. SynopsisWe introduce the atomic box correlation coefficient (aBCC), a machine learning-based metric for the resolution-standardized atom-level evaluation of ligand coordinate-density consistency in crystallographic structures. aBCC provides a common framework for assessing and communicating the local coordinate reliability between structural biologists and researchers in structure-based drug discovery.
Shtyrov, A.; Wilson, H.; Murshudov, G. N.
Show abstract
Damage to biological specimens by the electron beam is the fundamental resolution-limiting factor in cryoelectron microscopy (cryo-EM) single particle analysis. There is, however, currently no method to accurately infer fluence-dependent changes to the specimen structure during electron irradiation. We develop a Bayesian framework to fit a sequence of atomic models to a series of cryo-EM reconstructions produced at increasing fluence. In particular, our algorithm is able to infer the ensemble average position and atomic displacement parameter of every atom in the macromolecule as a function of fluence. Application of the algorithm to cryo-EM datasets shows that the molecule expands during imaging and identifies environment-dependent variations in beam-induced damage. We use our results to propose a stochastic process model of this phenomenon. We envisage that our method will lead to a better mechanistic understanding of radiation damage to biological specimens and may contribute to efforts to mitigate its effects.
Peretroukhin, V.; McLean, M.; Punjani, A.
Show abstract
The quality of single particle cryo-EM reconstructions can be severely degraded when an insufficient variety of 3D particle orientations is present in the image data, limiting downstream model building and interpretation. However, it is often difficult to ascertain whether or not a particular dataset suffers from such preferred orientation since the required orientation coverage depends on target geometry, alignment accuracy, and particle quality. To simplify diagnosis of preferred orientation, we present two complementary methods. First, the conical Fourier Shell Correlation Area Ratio (cFAR) compares the worst- and best-correlating conical regions of 3D Fourier space to quantify half-map anisotropy into a single, easily interpretable score ranging from zero to one. Second, Relative Signal, a companion to cFAR, directly relates signal content to viewing direction so that under-sampled views can be identified. We characterize our methods and compare them to existing anisotropy detection approaches on synthetic data and on 14 real datasets that span sundry molecular weights and structure types. Implementations of both cFAR and Relative Signal are included in CryoSPARC v4.5 and later versions.
Nguyen, N.; Pham, B.
Show abstract
Single-particle cryo-electron microscopy (cryo-EM) pose estimation is traditionally solved anew for each dataset, where iterative refinement is done from scratch while the estimator learns to store the molecule in its weights. In this work, we show that pose inference is a generalizable, specimen-agnostic operation when conditioned explicitly on a reference volume. We introduce ARCHER, an amortized contrastive classifier that models the pose posterior over a discrete rotation grid. Trained across a variety of protein structures, it operates zero-shot without retraining per structure. This transferability is grounded in Fourier-space information mechanics, where all specimen dependence is captured by the reference structure's power spectrum and spatial extent. ARCHER achieves a median angular error of 5.0{degrees} on 100 held-out test structures and 2.5{degrees} on experimental particles, matching dedicated estimators within 0.16[A] in 3D reconstruction. Crucially, downstream conformational signal is preserved. The leading conformational coordinate correlates at 0.97 with deposited benchmarks, faithfully reconstructing free-energy basins and mobile domains. These results overall demonstrate that cryo-EM pose estimation can be generalized across different structures.
Friedl, A.; Manst, D.
Show abstract
Background: Comparisons between independently predicted wild-type and missense-variant protein structures can generate mechanistic hypotheses, but small apparent differences may reflect model-selection variability rather than mutation-specific effects. Methods: Human mitochondrial DNA polymerase gamma (POLG; UniProt P54098) variants p.Arg627Gln (R627Q) and p.Trp748Ser (W748S) were evaluated using five AlphaFold2-PTM network-model outputs per condition generated with one random seed under matched ColabFold settings. Ten pairwise wild type comparisons at each site described between-network model-selection variability. Variant effects were summarized across five within-network wild-type-versus-variant comparisons using rotation-invariant local C-alpha pair distances and local displacement after global and local alignment. Because these comparison designs differ, the wild-type distribution was used as context rather than a mutation-effect null. Wild-type cryo-EM structure 9GGF was used for contact and interface mapping. Experimental A467T and G848S structures 9GGE and 9GGC provided contextual benchmarks. Results: R627Q measurements fell within the range of between-network wild-type differences: its median mean local pair-distance change was 0.170 angstrom, compared with a wild-type median of 0.170 angstrom, and its locally aligned displacement was 0.265 versus 0.248 angstrom. W748S showed higher median values (0.168 versus 0.132 angstrom for pair-distance change; 0.236 versus 0.182 angstrom for locally aligned displacement), but the ranges overlapped and the comparison-design asymmetry precluded a calibrated mutation-effect percentile. Experimental A467T and G848S comparisons produced local changes of similar magnitude. In 9GGF, R627 and W748 directly shared a local microenvironment, with a minimum heavy-atom distance of 3.53 angstrom. R627 also formed short polar-contact candidates with D629 and D743, whereas W748 occupied a hydrophobic packing environment containing Y622 and F750. Both sites were more than 18 angstrom from nucleic acid, more than 30 angstrom from POLG2, and more than 33 angstrom from PZL-A in a ligand-bound structure. Conclusions: Available AlphaFold2 comparisons do not establish a mutation-specific structural deformation for either variant. Experimental-structure mapping supports testable physicochemical hypotheses involving a shared R627-W748 microenvironment - loss of an arginine-centered polar network for R627Q and disruption of a buried aromatic environment for W748S - but not direct DNA, POLG2, or PZL-A contact mechanisms. Matched control substitutions and independent seeds are required to calibrate small mutation-associated structural deltas.
Klein, I.; Agam, G.; Irving, T.
Show abstract
X-ray fiber diffraction patterns exhibit four-fold symmetry that can be exploited, through folding and averaging, to improve signal-to-noise ratio. Accurate folding requires a precise sub-pixel estimate of the symmetry center and precise orientation of the meridional pattern axis to the fiber axis: small center or angular errors blur diffraction features, reduce layer-line sharpness, and introduce errors in spacing measurements. A pixel-level estimate is often too imprecise for this purpose, and detector gaps further complicate the alignment objective. We formulate the masked quadrant-folding problem, define a four-quadrant symmetry loss that consistently excludes invalid pixels, and evaluate several refinement strategies: hierarchical coarse-to-fine grid search; ECC-based rigid registration with global center/orientation correction fitting; ECC registration followed by local gradient refinement; and a hybrid that appends a local grid search on a cropped pattern. Direct gradient optimization from the rough QF alignment was found to be unreliable. Grid search provides a robust, interpretable baseline that directly minimizes the folding objective but is substantially slower than registration; ECC gives a fast near-correct alignment, and the hybrid closes the accuracy gap to brute-force search at a fraction of its runtime. On real datasets with calibration data, applying a calibration center with optimized rotation is effectively optimal. The hybrid center-refinement method has been integrated into the MuscleX package.
Grunewald, L.; Meszaros, P.; Westenhoff, S.
Show abstract
Time-resolved serial crystallography (TR-SX) has emerged as a powerful method for capturing ultrafast structural dynamics in proteins. TR-SX continues to produce remarkable studies, revealing previously unobserved transient states and providing deeper insights into processes such as drug targeting, DNA repair, and photosynthesis. However, extracting weak structural signals from noisy time-resolved datasets remains a major challenge. Robust computational methods are therefore required to isolate the signals associated with the underlying transient states. Importantly, this should be performed in reciprocal space to preserve compatibility with established downstream structure refinement workflows. Here, we introduce a framework for kinetic decomposition directly in reciprocal space that enables separation of kinetically distinct structural states. The method decomposes crystallographic data according to a predefined kinetic model, improving the recovery of weak transient signals and enhancing mechanistic interpretation from limited time-resolved datasets. We validate the framework using simulated data based on a previously published time-resolved crystallography study and demonstrate its application to a new TR-SX dataset comprising 17 time points. We show that the method separates the reciprocal space signatures of four intermediates by incorporating kinetic information from a predefined reaction model. This establishes a workflow for extracting kinetic states directly from time-resolved X-ray diffraction data that can be seamlessly integrated into existing crystallographic structure-determination pipelines.
Chua, E. Y. D.; Rahmani, H.; Zhen, J.; Eisenstein, F.; Song, Y. H.; Johnston, J. D.; Wang, H.; Alink, L. M.; Kopylov, M.; Ho, C.-M.; Grotjahn, D.; de Marco, A.
Show abstract
Visualizing macromolecules within their native cellular context by cryo-electron tomography (cryo-ET) is fundamentally limited by the trade-off between field of view and resolution: Capturing high-resolution information about biomolecules requires high magnification, which restricts the field of view and obscures the cellular context in which those biomolecules function. Collecting montage data by tiling the electron beam over the region of interest offers one solution, although traditional round electron beams cause excessive radiation damage across overlapping regions. We previously made electron beams square in shape, enabling montage collection with minimal overlap and thereby reducing excessive exposure and loss of high-resolution information. Here, we create a pipeline for collecting and processing montage cryo-ET data with square electron beams. We show that square beam montages retain high-resolution information by reconstructing virus-like particles to 3.5 [A] resolution using sub-tomogram averaging, and apply the workflow to imaging a glial cell and malaria parasite lamellae over fields of view up to 65 m2. We also provide a comprehensive protocol to make square beams accessible to the community.
Simpkin, A. J.; Johnson, E.; Rigden, D.
Show abstract
Motivation: The actual interface pTM score (actifpTM) is a modified version of the ipTM score that limits the calculation to only those residues at the interface. Whilst actifpTM provides an effective interface quality score, a limiting factor is that it makes use of the predicted aligned error (PAE) with probabilities, information that is generated during a ColabFold run, but not output by the package or other model prediction software. The consequent inability to generate actifpTM scores for the results of software such as AlphaFold 2 or AlphaFold 3 has limited its adoption. With reactifpTM we address this problem by providing a standalone tool that can be run on the standard outputs of most model prediction packages. Results: Using the same underlying principles as actifpTM, reactifpTM has been developed to use standard output files from model prediction software (a model and corresponding PAE) to perform an actifpTM-like calculation. ColabFold models were generated for a dataset of 1079 known interfaces in the PDB. A strong correlation was shown between actifpTM and reactifpTM for this dataset. Availability and implementation: reactifpTM is coded in Python. All scripts and associated documentation are available from https://github.com/hlasimpk/reactifptm or https://pypi.org/project/reactifptm.
Srivastava, V.; Mai, H.; Collins, M.; HOLTON, J. M.; Wall, M.; Wankowicz, S. A.
Show abstract
Ordered water molecules mediate many protein functions, including stability, ligand binding, and catalysis. Predicting their positions with sub-angstrom accuracy would support protein design, binding affinity prediction, and automated model building in X-ray crystallography and cryo-EM. However, water molecule prediction lags behind protein and other molecule structure predictions. Here, we introduce WaterFlow, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures. WaterFlow outperforms the existing state of the art at every precision level. We demonstrate that WaterFlow can accurately predict ground truth modeled water molecules, including those around protein-ligand interactions and on predicted structures. We also show that WaterFlow predictions fit well directly to experimental data, and therefore propose that it may be used for both prediction and modeling water molecules. This includes novel predictions that are often associated with positive electron difference density, meaning the model places water molecules at sites the original structure depositions omitted. We use this improved model to address the data constraint. By mapping the Pareto front of achievable accuracy of water molecule prediction, alongside analysis of different training data schemas, we quantified the trade-off between data quantity and data quality, demonstrating that the diversity of high-quality structures is limiting the possible results. Overall, WaterFlow predicts ordered water to serve as a solvent module for structure-based drug design and for water molecule placement during crystallographic refinement.
Monrroy, L.; Cardoch, S.; Westenhoff, S.
Show abstract
Solution X-ray scattering provides unique structural information on biomolecules under biological conditions, resolving conformational heterogeneity and time-resolved structural changes. The scattering profiles contain limited information, and interpretation largely relies on fitting candidate structures guided by priors. Direct reconstruction of electron density maps is desirable, but so far has been prevented by the difficulty of incorporating such prior knowledge. Here we propose XSSDense, a framework that couples a variational autoencoder trained on electron densities from predicted or simulated protein ensembles with a genetic algorithm to refine densities against scattering data. We validate XSSDense on synthetic data for crambin, recover the conformational heterogeneity of the unfolded state of Avena sativa light-oxygen-voltage sensing domain 2, resolve a de-novo density for the pre-unfolding state of the same protein, and provide a new structural description of the signalling-state ensemble of photoactive yellow protein. XSSDense enables structurally grounded electron density reconstructions that intrinsically capture conformational heterogeneity.
Tanino, H.; Tsujino, H.; Nakao, T.; Oie, C.; Makino, F.; Miyata, T.; Kasai, K.; Namba, K.; Inoue, T.
Show abstract
Human cytochrome P450 2C9 (CYP2C9) is a hepatic microsomal enzyme involved in the oxidative metabolism of clinically important drugs, but the structural organization of its oligomeric assemblies outside crystallographic packing environments remains poorly understood. Here, we report the cryo-EM structure of human CYP2C9 determined under aqueous, membrane-free conditions at 3.31 Angstrom resolution. The structure reveals a C2-symmetric hexameric assembly organized as a dimer of trimers. Individual protomers retain the conserved P450 fold and heme-binding architecture observed in previously reported crystal structures, indicating that assembly formation does not substantially perturb the catalytic core. The hexamer is stabilized by defined intra-trimer interfaces involving the N-terminal region and residues around Trp212 and Phe482, together with inter-trimer interfaces involving Leu71 and the 220-227 loop. These interfaces are distinct from the crystal packing contacts observed in CYP2C9 crystal structures, demonstrating that the assembly is not a simple recapitulation of crystallographic packing. Notably, the inter-trimer interface is located near the FG-loop-containing surface previously implicated in membrane association. This suggests that the observed hexamer may represent a membrane-free association of two trimers through membrane-related surfaces, whereas the trimeric arrangement itself may be compatible with membrane-associated organization. The structure therefore provides a framework for investigating how trimer formation, membrane interaction and local conformational changes in the FG-loop region may influence CYP2C9 function.
Gorelick, S.; Trepout, S.; Cleeve, P.; Boudes, M.; Kim, Y.; Ramm, G.
Show abstract
Preparing electron-transparent cryo-lamellae is inherently a serial, low-throughput process. During sample handling, milling, and transfer, cryo-fixed cells and their supporting films are subjected to mechanical forces as well as thermal stresses caused by temperature fluctuations. After milling, these extremely thin lamellae remain vulnerable to both mechanical and thermal stress, often leading to cracking or complete disintegration. Consequently, the loss of valuable lamellae is frequently an unavoidable aspect of working with such fragile specimens. In this work, we reconsider the conventional lamella geometry, which is typically a flat, thin cross-sectional slab. During milling, lamellae often become unintentionally bent, complicating the final polishing step required to achieve uniform thinning across their width. To address this limitation, we propose deliberately fabricating lamellae in a pre-bent configuration, i.e. specifically, adopting an arch-shaped profile instead of the traditional flat geometry. The arch shape is intrinsically more mechanically stable than a flat structure, thereby reducing lamella loss due to mechanical failure. Moreover, pre-bent milling patterns facilitate uniform thinning of bent lamellae, which is difficult to achieve using conventional flat milling approaches. In addition to the arch geometry, we investigate corrugated lamellae, characterised by a sinusoidal variation around the plane of a conventional flat lamella. Similarly to the arch shape, the corrugated design offers enhanced mechanical stability compared to traditional flat lamellae. We fabricated a series of test lamellae incorporating both arches and corrugations. High-resolution cryo-TEM imaging was performed to evaluate these structures, demonstrating that non-flat geometries do not compromise cryo-electron tomography performance. Furthermore, finite element method (FEM) simulations were conducted to provide insight into stress distributions within bent and corrugated lamellae.
Krupyanskii, Y. F.; Kovalenko, V.; Loiko, N.; Generalova, A.; Tereshkin, E.; Tereshkina, K.; Sokolova, O.; Peters, G.
Show abstract
This paper presents and critically reviews the results of original and some literature based experimental studies conducted by the authors last years on the structural organization of DNA in dormant (starvation stress), anabiotic dormant (4 HR treatment) E. coli cells, as well as the K12 {Delta}dps strain, which lacks the Dps protein (Dps null E. coli). The experimental data includes small-angle synchrotron radiation diffraction (SAXS) and transmission electron microscopy (TEM) data. Synchrotron radiation diffraction experiments on K12{Delta}dps cells allowed us to conclude that peaks at 44.3, 22.1, and 14.8 angstrom resolutions are associated exclusively with ordered DNA organization. Peaks at 44.3, 22.1, and 14.8 angstrom resolutions are also observed for samples of dormant (starvation stress) cells and anabiotically dormant cells. Therefore, this ordered DNA organization also applies to samples of dormant and anabiotically dormant cells. A model is proposed that considers the ordered DNA organization in the cell as a cholesteric liquid crystal. The powder diffraction pattern calculated based on this model is compared with experimental small angle X ray scattering (SAXS) data obtained on Dps-null cell samples. The model completely reproduces the key features of the experimental diffraction pattern from Dps-null cell samples. Accordingly, the cholesteric liquid crystal model corresponds to DNA packaging in dormant and anabiotically dormant cells. Cholesteric liquid crystal ordering should be further considered in all models of cellular DNA packaging. To address the question of which structural organization of DNA predominates in the cell: the cholesteric liquid crystal or nanocrystalline or whether they coexist and fully manifest themselves under different external conditions, it is necessary to utilize the latest methodological advances in structural analysis.
Smirnov, S. L.; Vugmeyster, L.; Stephenson, N.; McCarty, J.
Show abstract
Biophysics is a rapidly advancing field with an incredible breadth of topics. Thus, undergraduate biophysics instructors have to strategize and decide what topics they will cover in their courses. Educational institutions utilize a variety of biophysics textbooks. A common deficiency of each of the existing texts is that it serves well a given set of topics (theory, illustrations, practice problems) and leaves out other areas. A typical example includes good theory and problems for thermodynamics and kinetics while presenting molecular dynamics and various spectroscopic methods in a lacking or outdated way. The authors of this manuscript teach a capstone Biophysical Chemistry three-quarter series (Western Washington University/WWU, Bellingham, WA) which ideally should resonate with the general and major-specific courses the students take within their major at WWU. To achieve this goal and to enrich the traditional lecture-based delivery, the instructors have developed and brought together key pedagogical elements: purpose-built online textbook with a uniform structure of the academic content and practice problems, a study sample (oligopeptide) of biophysical significance with a growing set of experimental and computational data and student-centric in-class activities including computer labs. Our Biophysical series emphasizes concepts and methods of computational structural biology (Molecular Dynamics) and spectroscopic approaches (IR, UV and NMR). Here we describe the details of our integrative approach, summarize key outcomes and chart ways to advance the biophysical chemistry series further. Our textbook can be found through LibreText.
Baker, T. H.; Ohi, M. D.; Salmen, W.
Show abstract
Proteins and their associated complexes often adopt multiple conformations, with the transitions between these states playing a critical role in biological function. However, the resulting structural heterogeneity can be challenging to visualize and communicate, often requiring manual inspection and time-consuming annotation of biomolecular structures. To address this, we developed ResiRuler, a local, browser-based tool that uses inter-residue distance measurements to quickly quantify atomic displacements and map changes in internal geometry across ensembles of related protein structures. By converting structural differences into residue-pair distance changes, ResiRuler enables rapid identification of regions undergoing coordinated motion, local rearrangement, or large-scale conformational change. The resulting visualizations can be exported as scripts for PyMOL and ChimeraX, allowing users to explore conformational differences and generate publication-quality molecular figures in their preferred visualization environment. Using atomic models in Macromolecular Crystallographic Information File (mmCIF) file format, ResiRuler aligns multiple structures and measures structural variation across models facilitating visualization and presentation of these differences. This allows for rapid visualization of which regions of proteins change among ensembles of structures. The program is available for download at https://github.com/tbaker67/ResiRuler on macOS and Linux operating systems.
Zehnacker, S.; Caffarri, S.; Blanc, G.; Johnson, X.; Siponen, M.
Show abstract
RationaleRecent viral metagenomic studies have identified a plethora of enzyme-encoding genes in Phycodnaviridae viruses that are not strictly required for viral replication. These enzymes hold an unexpected metabolic potential during the infection process with their specific green algae host. As neither their role in the infection process nor the subcellular localization of these proteins has been experimentally characterized, comparative sequences, structural and biochemical in silico analyses can help generate functional and localization hypotheses. MethodsIn a recent viral metagenomic dataset, we identified a collection of viral homologs involved in bilin biosynthesis: heme oxygenase (vHMOX1) and Phycocyanobilin:Ferredoxin oxidoreductase (vPcyA). Viral and algal homologues were compared through sequence analyses and AlphaFold3 structural predictions. Predicted biochemical properties were analyzed for their compatibility with subcellular compartments. Active site architecture and putative substrate binding were compared between viral and algal proteins using AlphaFold3 and experimentally resolved structures. ResultsViral HMOX1 and PcyA sequences are truncated compared to algal homologs, lacking the N-terminal extension associated with chloroplast targeting. However biochemical properties, including isoelectric point and surface charge distribution, are compatible with localization in chloroplast stroma. Structural comparisons reveal modifications in the viral HMOX1 active site, including partial substrate reorientation and substitutions of key residues, consistent with modified heme-binding properties. In contrast, vPcyA models show no significant differences to their algal counterparts. ConclusionsActive site remodeling in vHMOX1 protein models suggests that these viral homologues may have evolved distinct heme-binding properties. Unlike vPcyA, vHMOX1 homologs appear to have diverged more substantially from their algal counterparts, potentially reflecting functional specialization in the viral infection context. One sentence summary of key findingsOur bioinformatic analyses expand the repertoire of auxiliary metabolic genes in Phycodnaviridae by identifying a conserved heme degradation pathway, non-canonical vHMOX1/PcyA targeting and structural rearrangements surrounding the catalytic sites of viral HMOX1.
Gall, L.; Shirgill, S.; Abbott, H.; Nieves, D. J.; Owen, D. M.
Show abstract
Quantitative analysis of single-molecule localisation microscopy (SMLM) data remains challenging because biologically diverse, well-annotated datasets are limited, whilst nanoscale protein organisation is heterogeneous and difficult to describe with hand-tuned metrics. We present SynthMLM, a framework that infers interpretable structural descriptors from experimental SMLM data and uses these descriptors to generate synthetic localisation datasets. We demonstrate SynthMLM by generating descriptor-matched synthetic datasets corresponding to diverse experimental SMLM datasets and evaluating their agreement with real data using descriptor-level and embedding-based measures. By enabling controlled generation of synthetic localisation data, SynthMLM provides a practical resource for benchmarking SMLM analysis methods, testing algorithm failure modes, and developing machine-learning workflows where large, labelled datasets are required.